MetaGeneTack: ab initio detection of frameshifts in metagenomic sequences
نویسندگان
چکیده
SUMMARY Frameshift (FS) prediction is important for analysis and biological interpretation of metagenomic sequences. Since a genomic context of a short metagenomic sequence is rarely known, there is not enough data available to estimate parameters of species-specific statistical models of protein-coding and non-coding regions. The challenge of ab initio FS detection is, therefore, two fold: (i) to find a way to infer necessary model parameters and (ii) to identify positions of frameshifts (if any). Here we describe a new tool, MetaGeneTack, which uses a heuristic method to estimate parameters of sequence models used in the FS detection algorithm. It is shown on multiple test sets that the MetaGeneTack FS detection performance is comparable or better than the one of earlier developed program FragGeneScan. AVAILABILITY AND IMPLEMENTATION MetaGeneTack is available as a web server at http://exon.gatech.edu/GeneTack/cgi/metagenetack.cgi. Academic users can download a standalone version of the program from http://exon.gatech.edu/license_download.cgi.
منابع مشابه
GeneTack database: genes with frameshifts in prokaryotic genomes and eukaryotic mRNA sequences
Database annotations of prokaryotic genomes and eukaryotic mRNA sequences pay relatively low attention to frame transitions that disrupt protein-coding genes. Frame transitions (frameshifts) could be caused by sequencing errors or indel mutations inside protein-coding regions. Other observed frameshifts are related to recoding events (that evolved to control expression of some genes). Earlier, ...
متن کاملFrameshift alignment: statistics and post-genomic applications
MOTIVATION The alignment of DNA sequences to proteins, allowing for frameshifts, is a classic method in sequence analysis. It can help identify pseudogenes (which accumulate mutations), analyze raw DNA and RNA sequence data (which may have frameshift sequencing errors), investigate ribosomal frameshifts, etc. Often, however, only ad hoc approximations or simulations are available to provide the...
متن کاملGenetack: frameshift Identification in protein-Coding Sequences by the Viterbi Algorithm
We describe a new program for ab initio frameshift detection in protein-coding nucleotide sequences. The task is to distinguish the same strand overlapping ORFs that occur in the sequence due to a presence of a frameshifted gene from the same strand overlapping ORFs that encompass true overlapping or adjacent genes. The GeneTack program uses a hidden Markov model (HMM) of genomic sequence with ...
متن کاملBlastXtract2: Improving early exploration of (meta) genomes
UNLABELLED To manage and intelligently mine the avalanche of genomic sequences intuitive and user-friendly graphical interfaces are required. Here we present BlastXtract2 which exclusively facilitates early exploration of un-annotated genomic and metagenomic sequences. Various formats of translated searches, including the commonly used BlastX, of multiple sequences against multiple protein data...
متن کاملInvestigation of Solvent Effect on CUA Codon Mutation: NMR Shielding Study
P53 is one of the gene that has important role in human cell cycle and in the human cancers too.Models of codon substitution make it possible to separate mutational biases in the DNA fromselective constraints on the protein, and offer a great advantage over amino acid models forunderstanding the evolutionary process of proteins and protein-coding DNA sequences. In thiswork, we investigated abou...
متن کامل